Papers with annotation tasks

29 papers
Geo-Cultural Representation and Inclusion in Language Technologies (2024.lrec-tutorials)

Copied to clipboard

Challenge: audi et al.: training and evaluation of language models rely on semi-structured data that is annotated by humans . e-learning tools do not integrate rich and diverse community perspectives into language technologies .
Approach: They will examine how different socio-cultural perspectives influence what is taken as ground truth by models.
Outcome: This tutorial examines how different socio-cultural perspectives influence representations of global concepts.
AnnIE: An Annotation Platform for Constructing Complete Open Information Extraction Benchmark (2022.acl-demo)

Copied to clipboard

Challenge: Open Information Extraction (OIE) is the task of extracting facts from sentences in the form of relations and their corresponding arguments in schema-free manner.
Approach: They propose an interactive annotation platform that facilitates annotating complete facts from input sentences.
Outcome: The proposed platform facilitates such challenging annotation tasks and supports creation of fact-oriented OIE evaluation benchmarks.
HUMAN: Hierarchical Universal Modular ANnotator (2020.emnlp-demos)

Copied to clipboard

Challenge: HUMAN is a web-based annotation tool that covers a variety of annotation tasks on textual and image data.
Approach: They propose a web-based annotation tool that covers a variety of annotation tasks on textual and image data.
Outcome: HUMAN covers a variety of annotation tasks on textual and image data and uses an internal deterministic state machine to chain different tasks in an interdependent manner.
Integrating INCEpTION into larger annotation processes (2024.emnlp-demo)

Copied to clipboard

Challenge: Annotation tools are increasingly only steps in a larger process into which they need to be integrated.
Approach: They propose to adapt INCEpTION, a semantic annotation platform that offers intelligent assistance and knowledge management.
Outcome: The proposed platform offers a range of APIs and can interact with external services such as authorization services, crowdsourcing platforms, terminology services or machine learning services.
ActiveAnno: General-Purpose Document-Level Annotation Tool with Active Learning Integration (2021.naacl-demos)

Copied to clipboard

Challenge: Existing tools for document-level annotation lack document-based quality and flexibility.
Approach: a new annotation tool is being developed for industry and research use cases . a configurable user interface and a RESTful API are included . authors propose to use ACTIVEANNO as default for document-level annotation .
Outcome: ACTIVEANNO is an annotation tool for industry and research use cases.
WARP-Text: a Web-Based Tool for Annotating Relationships between Pairs of Texts (C18-2)

Copied to clipboard

Challenge: Existing tools for annotating pairs of texts do not support detailed pairwise annotation.
Approach: They present an open-source web-based tool for annotating relationships between pairs of texts . they propose to use WARP-Text to create multi-layer annotations and custom definitions .
Outcome: The proposed tool can be used by project managers and annotators.
POTATO: The Portable Text Annotation Tool (2022.emnlp-demos)

Copied to clipboard

Challenge: POTATO is a free, fully open-sourced annotation system that supports labeling many types of text and multimodal data.
Approach: They propose to use POTATO to design and deploy complex annotation tasks.
Outcome: The proposed annotation system improves labeling speed and productivity over two tasks.
QSTN: A Modular Framework for Robust Questionnaire Inference with Large Language Models (2026.eacl-demo)

Copied to clipboard

Challenge: Questionnaire-like prompts have become an important format to probe, assess, and utilize large language models (LLMs)
Approach: They propose an open-source Python framework for generating responses from questionnaire-style prompts to support in-silico surveys and annotation tasks with large language models (LLMs).
Outcome: The proposed framework can be used to generate responses from questionnaire-style prompts and to perform annotations on large language models.
Potato 2.0: A Comprehensive Annotation Platform with AI-in-the-Loop Support (2026.acl-demo)

Copied to clipboard

Challenge: Annotated data is still central to NLP and Generative AI, yet the demands on annotation have grown in both scale and complexity.
Approach: They introduce Potato 2.0, an open source annotation platform for easy deployment and customization.
Outcome: The new version of potato supports 39 different types of annotation tasks and multiple AI-assistance features.
Design Choices for Crowdsourcing Implicit Discourse Relations: Revealing the Biases Introduced by Task Design (2023.tacl-1)

Copied to clipboard

Challenge: Disagreement in natural language annotation has been studied from a perspective of biases introduced by the annotators and the annotation frameworks.
Approach: They propose to analyze task design bias in crowdsourced annotations where lay annotators are used to elicit interpretations.
Outcome: The proposed methods can push annotators towards certain relations and some discourse relation senses can be better elicited with one or the other approach.
Corpus Considerations for Annotator Modeling and Scaling (2024.naacl-long)

Copied to clipboard

Challenge: Recent trends in natural language processing and annotation tasks emphasize individual perspectives . annotator models that rely on a single ground truth may disregard valuable minority perspectives omissions .
Approach: They propose a composite embedding approach to investigate annotator modeling techniques . they show that the commonly used user token model consistently outperforms more complex models .
Outcome: The proposed model outperforms more complex models on a given dataset.
“All that Glitters”: Techniques for Evaluations with Unreliable Model and Human Annotations (2025.findings-naacl)

Copied to clipboard

Challenge: Using standard metrics in the presence of poor labels masks label and model quality . evaluation techniques accounting for unreliable labels reveal important flaws, including spurious correlations and nonrandom racial biases .
Approach: They analyze human labels, GPT model ratings, and transformer encoder model ratings . they show that standard metrics in the presence of poor labels mask label and model quality .
Outcome: The proposed methods mask label and model quality even in the presence of poor models.
Toward Annotator Group Bias in Crowdsourcing (2022.acl-long)

Copied to clipboard

Challenge: Annotator group bias is a common problem in crowdsourcing, but is often overlooked .
Approach: They propose a probabilistic framework to capture annotator group bias using an extended Expectation Maximization algorithm.
Outcome: The proposed model can model annotator group bias over competitive datasets and demonstrate that it is effective over multiple datasets.
Fine-Grained Error Analysis and Fair Evaluation of Labeled Spans (2022.lrec-1)

Copied to clipboard

Challenge: Annotations with incorrect label or boundaries count as two errors instead of one, despite being closer to the target annotation than false positives or false negatives.
Approach: They propose an algorithm for error identification in flat and multi-level annotations and propose a procedure for calculating meaningful precision, recall, and F1-scores based on the more fine-grained error types.
Outcome: The proposed procedure prevents double penalties and allows for a more detailed error analysis, providing more insight into the actual weaknesses of a system.
Text Annotation Graphs: Annotating Complex Natural Language Phenomena (L18-1)

Copied to clipboard

Challenge: Text Annotation Graphs is a web-based tool for annotating text . it provides functionality for representing complex relationships between words and word phrases .
Approach: They introduce a web-based tool for annotating text, Text Annotation Graphs, or TAG . it provides functionality for representing complex relationships between words and word phrases .
Outcome: The proposed software can represent complex relationships between words and words . it can also be used to find similar structures within the current document or external annotated documents.
TurkingBench: A Challenge Benchmark for Web Agents (2025.naacl-long)

Copied to clipboard

Challenge: TurkingBench is a benchmark consisting of tasks presented as web pages with textual instructions and multi-modal contexts.
Approach: They propose to use HTML pages to perform various annotation tasks on crowdsourcing platforms.
Outcome: The proposed model outperforms other models on the TurkingBench benchmark.
Fluid Annotation: A Granularity-aware Annotation Tool for Chinese Word Fluidity (L18-1)

Copied to clipboard

Challenge: Using word segmentation, we propose a wordhood annotation framework for Chinese language . word segmentations have been used for years in preprocessing NLP tasks for languages without explicit word delimiter.
Approach: They propose a word-granularity-aware annotation framework for Chinese language . they argue that word segmentation is fluid in nature and that it rearranges the boundary of word segmentations and linguistic annotation.
Outcome: The proposed framework rearranges the boundary between word segmentation and linguistic annotation and supports flexible annotation tasks for various linguistic and affective phenomena.
Characterizing Human and Zero-Shot GPT-3.5 Object-Similarity Judgments (2024.findings-naacl)

Copied to clipboard

Challenge: Recent advances in large language models have yielded few-shot, human-comparable performance on a range of tasks, but studies of LLM annotation accuracy and behavior are sparse.
Approach: They characterize OpenAI’s GPT-3.5’s judgment on a behavioral task for implicit object categorization and give similarities and differences between them.
Outcome: The proposed model augments human responses with LLMs for domains where data is sparse or compute resources are low.
Lessons Learned from a Citizen Science Project for Natural Language Processing (2023.eacl-main)

Copied to clipboard

Challenge: Annotations are expensive and difficult to obtain, which is why many NLP systems outsource their work to paid crowdworkers.
Approach: They propose to use Citizen Science to re-annotate parts of a pre-existing crowdsourced dataset to gain high-quality annotations.
Outcome: The proposed approach yields high-quality annotations and motivated volunteers, but requires consideration of scalability, participation over time, and legal and ethical issues.
Simple Semantic Annotation and Situation Frames: Two Approaches to Basic Text Understanding in LORELEI (L18-1)

Copied to clipboard

Challenge: Existing annotations for low resource languages are under-resourced for human language technology, but lack of resources does not correlate with lack of need for such technologies.
Approach: They propose two types of semantic annotation for the DARPA Low Resource Languages for Emerging Incidents program: Simple Semantic Annotation (SSA) and Situation Frames (SF).
Outcome: The proposed approaches are aimed at labeling basic semantic information relevant to humanitarian aid and disaster relief scenarios.
ARAIDA: Analogical Reasoning-Augmented Interactive Data Annotation (2024.acl-long)

Copied to clipboard

Challenge: Empirical studies demonstrate that Araida reduces human correction labor by 11.02% compared to vanilla interactive data annotation methods.
Approach: They propose an analogical reasoning-based approach that enhances automatic annotation accuracy in the interactive data annotation setting and reduces the need for human corrections.
Outcome: Empirical studies show that Araida reduces human correction labor by 11.02% compared to vanilla interactive data annotation methods.
Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It’s Best to Relate Perspectives! (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to subjectivity in natural language processing are subjective . authors argue that disagreement should not be regarded as a problem .
Approach: They propose to account for subjective perspectives of individuals and objective concepts that build a common ground between annotators.
Outcome: The proposed architectures increase the averaged annotator-individual F1-scores up to 43% over a majority-label model.
Discourse Analysis via Questions and Answers: Parsing Dependency Structures of Questions Under Discussion (2023.findings-acl)

Copied to clipboard

Challenge: Existing discourse formalisms require large taxonomies of discourse relations to be accurate.
Approach: They propose a linguistic framework for discourse analysis using questions under discussion . they propose qUD parser that derives a dependency structure of questions over full documents .
Outcome: The proposed model is trained on a large, crowdsourced question-answering dataset.
TabEmb: Joint Semantic-Structure Embedding for Table Annotation (2026.acl-long)

Copied to clipboard

Challenge: Existing tables learn by linearizing the 2D table into a 1D token sequence and encoding it with pretrained language models (PLMs) such as BERT, but this leads to limited semantic quality and weaker generalization to unseen or rare values compared to modern LLMs.
Approach: They propose a table annotation module called TabEmb which decouples semantic encoding from structural modeling by creating semantically rich embeddings for each column.
Outcome: Experiments show that TabEmb outperforms baselines on different table annotation tasks.
A Needle in a Haystack: An Analysis of High-Agreement Workers on MTurk for Summarization (2023.acl-long)

Copied to clipboard

Challenge: Using crowdsourcing, it is difficult to obtain high-quality annotations for difficult tasks.
Approach: They propose a recruitment pipeline to recruit high-quality Amazon Mechanical Turk workers . they filter out subpar workers before they carry out the evaluations .
Outcome: The proposed method can filter out subpar workers before they carry out evaluations and obtain high-agreement annotations with similar constraints on resources.
A Web-based Collaborative Annotation and Consolidation Tool (2020.lrec-1)

Copied to clipboard

Challenge: Annotation tools have a rigid structure, closed back-end and front-end, and are built in a non-user-friendly way rendering them unusable for a large cohort.
Approach: They propose a web-based collaborative annotation and consolidation tool (AWOCATo) that supports varied textual formats and allows users to easily adapt to the annotation task.
Outcome: AWOCATo supports a range of tasks and domains, filling the gap left by the lack of tools that can be used by people with and without programming knowledge.
Annotating the Annotators: Analysis, Insights and Modelling from an Annotation Campaign on Persuasion Techniques Detection (2025.findings-acl)

Copied to clipboard

Challenge: Existing annotation campaigns based on heuristic guidelines have not been thoroughly discussed.
Approach: They propose a probabilistic model for optimizing intervention scheduling to reduce the cost of an expert oversight in annotation tasks.
Outcome: The proposed model advocates for an expert oversight in annotation tasks and periodic quality audits to reduce costs.
T5Score: A Methodology for Automatically Assessing the Quality of LLM Generated Multi-Document Topic Sets (2025.findings-acl)

Copied to clipboard

Challenge: Existing evaluation methods for Multi-Document Topic Extraction are not designed for LLMs and result in low inter-annotator agreement scores.
Approach: They propose an evaluation methodology that decomposes the quality of a topic set into quantifiable aspects, measurable through easy-to-perform annotation tasks.
Outcome: The proposed evaluation methodology decomposes the quality of a topic set into quantifiable aspects, measurable through easy-to-perform annotation tasks.
R1-RE: Cross-Domain Relation Extraction with RLVR (2026.acl-long)

Copied to clipboard

Challenge: Relation extraction (RE) is a core task in natural language processing.
Approach: They propose a supervised learning task for relation extraction (RE) based on annotation guidelines.
Outcome: The proposed model achieves an average OOD accuracy of 70%, on par with leading proprietary models such as GPT-4o.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations